Skip to content

feat(events): add chain_reorg and safe_target chain events - #533

Closed
MegaRedHand wants to merge 1 commit into
mainfrom
feat/events-reorg-safe-target
Closed

feat(events): add chain_reorg and safe_target chain events#533
MegaRedHand wants to merge 1 commit into
mainfrom
feat/events-reorg-safe-target

Conversation

@MegaRedHand

@MegaRedHand MegaRedHand commented Jul 21, 2026

Copy link
Copy Markdown
Collaborator

Continues the chain-event pub-sub series (#516 / #517 / #518, all merged). Adds the two cheap follow-up topics, both derived from the existing actor-layer snapshot diff.

Events added

Topic Payload Emitted when
chain_reorg { slot, depth, old_head_block, old_head_state, new_head_block, new_head_state } Fork choice switches to a head off the old head's chain (same recency gate as head)
safe_target { slot, block } The interval-3 fork-choice safe attestation target advances

Notes

  • chain_reorg reuses store::reorg_depth, widened to pub(crate) so the actor diff and block processing share one definition of "reorg". Recency-gated like head (HEAD_EVENT_RECENCY_SLOTS) and emitted just before it, matching beacon ordering; payload mirrors the beacon chain_reorg shape minus epoch/execution_optimistic.
  • safe_target is an ethlambda extension with no beacon analog; ungated since it coalesces to the latest value (the only one that matters).

Testing

cargo fmt, cargo clippy -p ethlambda-blockchain (clean), blockchain lib tests pass, including two new unit tests: chain_event_diff_emits_chain_reorg_before_head and chain_event_diff_emits_safe_target.


Independent PR, based off main. One of three sibling PRs continuing the series — alongside #534 (attestation/aggregate) and #535 (block_gossip), each independently based off main. They touch overlapping regions of events.rs / docs/rpc.md, so whichever two merge later will each need a small conflict rebase. Opened as draft.

Extend the /lean/v0/events stream with two cheap follow-up topics
derived from the actor-layer snapshot diff.

chain_reorg: fork choice switched to a head off the old head's chain.
Reuses store::reorg_depth (widened to pub(crate) so the actor diff and
block processing share one definition of "reorg"); recency-gated like
head and emitted just before it. Mirrors the beacon chain_reorg payload
minus epoch/execution_optimistic.

safe_target: the interval-3 fork-choice safe attestation target
advanced. Ungated since it coalesces to the latest value, the only one
that matters; an ethlambda extension with no beacon analog.
MegaRedHand added a commit that referenced this pull request Jul 24, 2026
Continues the chain-event pub-sub series (#516 / #517 / #518, all
merged). Adds the beacon `block_gossip` analog.

## Event added

| Topic | Payload | Emitted when |
|-------|---------|--------------|
| `block_gossip` | `{ slot, block }` | A block is seen on the network,
before import |

## Notes

- Emitted from the blockchain actor's `NewBlock` handler, **not** a
second P2P-side publisher, so the actor stays the sole event publisher
and the write flow stays one-directional (a core requirement of the
series' design).
- Fires for both gossip'd and by-root-fetched blocks, since both
re-enter through `NewBlock`. Import may pend the block while its parent
chain is fetched, so this precedes import.
- Ungated like `block` (not recency-gated), so subscribers can watch
sync progress. Low-rate (one per block), so the payload is built
unconditionally behind `emit`'s no-subscriber guard.

## Testing

`cargo fmt`, `cargo clippy -p ethlambda-blockchain -p ethlambda-rpc`
(clean), blockchain lib + rpc events tests pass.

---
**Independent PR, based off `main`.** One of three sibling PRs
continuing the series — alongside #533 (`chain_reorg`/`safe_target`) and
#534 (`attestation`/`aggregate`), each independently based off `main`.
They touch overlapping regions of `events.rs` / `docs/rpc.md`, so
whichever two merge later will each need a small conflict rebase. Opened
as draft.
@MegaRedHand

Copy link
Copy Markdown
Collaborator Author

Closing this until the events become relevant

@MegaRedHand
MegaRedHand deleted the feat/events-reorg-safe-target branch July 24, 2026 20:29
MegaRedHand added a commit that referenced this pull request Jul 24, 2026
Continues the chain-event pub-sub series (#516 / #517 / #518, all
merged). Adds the two high-rate topics.

## Events added

| Topic | Payload | Emitted when |
|-------|---------|--------------|
| `attestation` | `{ validator_id, data: AttestationData }` | A single
validator vote passes gossip validation (signature omitted) |
| `aggregate` | `{ participants: [u64], data: AttestationData }` | A
committee-signature aggregate is produced locally or accepted from
gossip (proof omitted) |

Both fire only for messages the store **accepted** (data + signature
validation), so subscribers see the same votes fork choice does. A node
never receives its own aggregate back over gossip, so `aggregate` fires
at two sites: the locally produced `AggregateProduced` path and the
gossip-received path.

## Design points

- The ~3 KB XMSS signature and the SNARK proof bytes are deliberately
**omitted** — too heavy for a high-rate stream.
- The shared broadcast channel capacity is bumped **256 → 8192**. All
topics share one ring buffer, so a subscriber's tolerable stall is
`capacity / total_event_rate` (dominated by the attestation rate), not
per-topic. A per-topic split behind the `EventBus` facade remains the
escape hatch if real devnet rates show the shared window biting — see
the note in `docs/rpc.md`.
- Emission uses the plain, no-subscriber-guarded `emit`. Consequence:
the per-vote `AttestationData` is cloned even when nobody is subscribed
(the guard only skips the send). If that actor-hot-path cost proves
material, a payload-deferral helper can be added in a follow-up.

## Testing

`cargo fmt`, `cargo clippy -p ethlambda-blockchain -p ethlambda-rpc`
(clean), blockchain + rpc lib tests pass, including
`events_streams_attestation_with_nested_data` (verifies the untagged
`data:` nesting over the wire).

---
**Independent PR, based off `main`.** One of three sibling PRs
continuing the series — alongside #533 (`chain_reorg`/`safe_target`) and
#535 (`block_gossip`), each independently based off `main`. They touch
overlapping regions of `events.rs` / `docs/rpc.md`, so whichever two
merge later will each need a small conflict rebase.
MegaRedHand added a commit that referenced this pull request Jul 29, 2026
## Summary

Adds `tooling/event-monitor/`, a standalone live arrival-time dashboard
for
lean-consensus (ethlambda) nodes. It dials the `GET /lean/v0/events` SSE
stream
of several nodes, timestamps each event on a single collector clock, and
serves
a browser dashboard visualizing:

- a **rolling beeswarm** of each event's arrival offset *within the
slot*, one lane per node, and
- a **propagation-delta beeswarm**: for a given block / aggregate, how
long after the *first* node each other node saw it.

The rolling window is adjustable live from the header, and a fresh page
load
backfills recent history from the collector so it is never blank.

## Design

- **Standalone Cargo workspace** under `tooling/event-monitor/` (empty
`[workspace]` table); it is not a member of the parent `ethlambda`
workspace and does not affect the main build.
- **Zero dependency on any ethlambda crate** — it only speaks the
documented SSE/HTTP wire shape. `CONTRACT.md` is the frozen interface
between the Rust/axum backend (`src/`) and the vanilla-JS frontend
(`web/`, no build step).
- A single collector clock makes propagation deltas skew-free across
nodes.
- `?demo=1` runs the frontend fully offline with synthetic data.

This PR contains the tooling plus the CI wiring needed to actually check
it. The
RPC / event-bus changes that make a node emit these events shipped
separately in
the chain-events series (below).

## Required ethlambda API functionality

All dependencies are **satisfied on `main`** — the default dashboard
view works
against a current node with no extra PRs.

| Capability | Endpoint / topic(s) | Status |
|---|---|---|
| SSE events stream | `GET /lean/v0/events` | ✅ #517 |
| Topic filtering | `?topics=<csv>` | ✅ #518 |
| Slot-geometry bootstrap | `GET /lean/v0/genesis`, `GET
/lean/v0/config/spec` | ✅ |
| Base topics | `block`, `head`, `justified_checkpoint`,
`finalized_checkpoint` | ✅ #516 |
| Attestation / aggregate topics | `attestation`, `aggregate` | ✅ #534 |
| Block-gossip topic | `block_gossip` | ✅ #535 |

The collector accepts `block`, `block_gossip`, `head`,
`justified_checkpoint`,
`finalized_checkpoint`, `attestation` and `aggregate`; any other topic
is logged
and skipped. `safe_target` / `chain_reorg` (#533, closed) were removed
in
4eeafff — both panels' topic filters discarded them, and `chain_reorg`
is a
point-in-time event rather than a per-node arrival race, so it does not
belong on
a beeswarm.

Both panels default to `block` / `attestation` / `aggregate`.

## Review fixes

Three defects found in review, each with a regression test:

- **Reconnect backoff never reset after a healthy session.** The reset
was keyed on a clean end-of-stream, but a node restart surfaces as a
stream *error*, so every restart ratcheted the delay one step
permanently; after ~7 of them a healthy node reported `down` with 10s
reconnects. Now keyed on session duration (survived ≥1 heartbeat). This
also closes the inverse case: a peer that accepts the request and
instantly closes the stream used to be retried at `INITIAL_BACKOFF`
forever.
- **One bogus slot blanked the dashboard.** Both the history ring and
the rolling window key retention off the highest slot seen, and that
watermark only moves up, so a single event from a node on a different
genesis aged out every real event until a restart. Events more than
`MAX_FUTURE_SLOTS` ahead of the collector's own slot are now dropped,
warning once per node per connection. Only the future side is bounded;
old slots are legitimate (`finalized_checkpoint` trails head) and cannot
move the watermark.
- **Stale status clobbered fresh status.** The frontend applies `status`
immediately but buffers `chain` during backfill, so the `/api/history`
snapshot could overwrite a newer live status. Live status now wins.

Also from that review:

- **Slot geometry is re-resolved every 60s** and republished when it
changes, dropping the retained history and slot watermark, so a
regenerated genesis no longer silently corrupts every subsequent
`offset_ms`. Collectors re-read per frame and `/api/meta` per request;
an already-open tab needs a reload to pick up new geometry.
- **The arrival axis now spans two slots at two scales**: the first at
full resolution (90.9% of the width at 5 intervals), the next compressed
into a tinted band half a first-slot interval wide, saturating beyond
that. A block spilling past its slot boundary is the failure mode worth
seeing, and it used to be indistinguishable from one landing exactly on
the boundary.
- **`offset_ms` includes the collector↔node round trip.** One clock is
what makes propagation deltas skew-free, but nodes reached over
different links carry a systematic offset that reads as lag. Now
documented in `README.md` and `CONTRACT.md §2`.
- History events are `Arc`-shared, so `publish_chain` no longer
deep-copies per event and `/api/history` no longer clones up to
`HISTORY_MAX_EVENTS` events while holding the mutex.
- `topics = []` / `nodes = []` are rejected at load rather than becoming
an opaque retry loop against a 400.

Known and deliberate, left as follow-ups: the collector cannot see
upstream
`: error - dropped N messages` comments (`eventsource-stream` swallows
them), so
a lagging subscription silently under-samples; both canvases still
repaint
unconditionally at 60fps; and `bootstrap` has no retry, so the monitor
exits if
started before its nodes are listening.

## CI

`tooling/event-monitor` declares its own `[workspace]`, so `cargo fmt
--all`,
`cargo check --workspace`, `cargo clippy --workspace`, `make lint` and
`make test`
all stop at the root workspace members and never reached it — it landed
with
nothing verifying it. The `Lint` job now also runs `cargo fmt --check`,
`cargo clippy --all-targets -- -D warnings` and `cargo test` under
`working-directory: tooling/event-monitor`, and `rust-cache` lists it as
a second
workspace so its target dir is cached and its `Cargo.lock` feeds the
cache key.

## Testing

- `cargo test` in `tooling/event-monitor/`: **41 passed** (39 lib + 2
integration), 0 failed.
- `cargo fmt --check` and `cargo clippy --all-targets -- -D warnings`:
clean.
- `?demo=1` offline synthetic mode renders both panels with no live
nodes.
- Verified end-to-end against a live local devnet (blocks land ~0.1–0.2
s into the slot, attestations ~0.8 s = interval 1).
- Two-slot axis geometry checked numerically: the overflow band is
exactly 0.5 × one first-slot interval, the mapping is monotonic, and the
endpoints land on the margins.
- Geometry refresh checked against a stub node: rewriting its
`genesis_time` mid-run is picked up within one refresh interval (`WARN
slot geometry changed`), `/api/meta` then serves the new epoch, and the
history ring holds the restarted chain's low slots with in-slot offsets
— which only works because the watermark is reset, otherwise they would
prune as older than the previous high-water mark.

## Run

```bash
cd tooling/event-monitor
cp config.example.toml config.toml
$EDITOR config.toml                 # list your nodes' RPC URLs + topics
cargo run --release -- --config config.toml
# open the `listen` address (default http://127.0.0.1:8080)
```
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant